A wavelet-based parameterization for speech/music discrimination
نویسندگان
چکیده
This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain. Résumé This paper addresses the problem of parameterization for speech/music discrimination. The current successful parameterization based on cepstral coefficients uses the Fourier transformation (FT), which is well adapted for stationary signals. In order to take into account the non stationarity of music/speech signals, this work proposes to study wavelet-based signal decomposition instead of FT. Three wavelet families and several numbers of vanishing moments have been evaluated. Different types of energy, calculated for each frequency band obtained from wavelet decomposition , are studied. Static, dynamic and long-term parameters were evaluated. The proposed parameterization are integrated into two class/non-class classifiers: one for speech/non-speech, one for music/non-music. Different experiments on realistic corpora , including different styles of speech and music (Broadcast News, Entertainment, Scheirer), illustrate the performance of the proposed parameterization, especially for music/non-music discrimination. Our parameterization yielded a significant reduction of the error rate. More than 30% relative improvement was obtained for the envisaged tasks compared to MFCC parameterization.
منابع مشابه
A wavelet-based parameterization for speech/music segmentation
The problem of speech/music discrimination is a challenging research problem which significantly impacts Automatic Speech Recognition (ASR) performance. This paper proposes new features for the Speech/Music discrimination task. We propose to use a decomposition of the audio signal based on wavelets, which allows a good analysis of non stationary signal like speech or music. We compute different...
متن کاملA new approach for audio classification and segmentation using Gabor wavelets and Fisher linear discriminator
Rapid increase in the amount of audio data demands an efficient method to automatically segment or classify audio stream based on its content. In this paper, based on the Gabor wavelet features, an audio classification and segmentation method is proposed. This method will first divide an audio stream into clips, each of which contains one-second audio information. Then, each clip is classified ...
متن کاملSpeech/Music Classification using wavelet based Feature Extraction Techniques
Audio classification serves as the fundamental step towards the rapid growth in audio data volume. Due to the increasing size of the multimedia sources speech and music classification is one of the most important issues for multimedia information retrieval. In this work a speech/music discrimination system is developed which utilizes the Discrete Wavelet Transform (DWT) as the acoustic feature....
متن کاملSpeaker Verification Reinforced by Objective Wavelet Packets-based Speech Parameterization
In attempt to enhance the discrimination ability of present speaker verification systems, we study alternative ways for parameterization of human voice. Utilizing the flexibility provided by wavelet packet analysis, and rendering an account of the most recent psycho-acoustical studies, the authors investigate in an objective way the relative importance of constituent disjoint frequency subbands...
متن کاملWavelet Parameterization for Speech Recognition
Typical parameterization schemes utilize linear prediction or melscaled filter-banks, which are classic windowed DFT based methods. In this paper a new optimized adaptive wavelet parameterization scheme is presented. A novel extension of the Best Basis algorithm is used on wavelet-packet cosine transform (WPCT) instead of typical filter bank. Obtained features are tested using Polish language H...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید
ثبت ناماگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید
ورودعنوان ژورنال:
- Computer Speech & Language
دوره 24 شماره
صفحات -
تاریخ انتشار 2010